Papers with noisy documents

6 papers
MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation (2025.acl-long)

Copied to clipboard

Challenge: Existing RAG systems struggle with the quality of retrieval documents, causing performance degradation and reducing performance.
Approach: They propose a training-free RAG framework that leverages multiple LLM agents to collaboratively filter and score retrieved documents.
Outcome: The proposed framework outperforms existing RAG frameworks in QA benchmarks and shows superior answer consistency and answer accuracy over baseline methods.
Typos that Broke the RAG’s Back: Genetic Attack on RAG Pipeline by Simulating Documents in the Wild via Low-level Perturbations (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on the robustness of Large Language Models (LLMs) overlook the interconnected relationships between RAG components or the potential threats prevalent in real-world databases, such as minor textual errors.
Approach: They propose a novel attack method that exploits vulnerabilities in RAG components and tests its robustness against noisy documents.
Outcome: The proposed method devastates the performance of each component and their synergy, and significantly devases the performance.
All That Glitters is Not Gold: Improving Robust Retrieval-Augmented Language Models with Fact-Centric Preference Alignment (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to learn adaptive retrieval for noisy documents lack prior filtering and may lead to the loss of crucial information.
Approach: They propose a method to improve retrieval performance without prior filtering . they use LLMs self-generated synthetic data as training data without manual annotation .
Outcome: The proposed method performs positive document mining based on factual consistency and uses LLMs self-generated synthetic data as training data without manual annotation.
Bootstrapped Pre-training with Dynamic Identifier Prediction for Generative Retrieval (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for document retrieval rely on static document identifiers . experimental results show that generative retrieval is outperforms dense retrieval in document retrievals.
Approach: They propose a bootstrapped pre-training method that dynamically adjusts document identifiers during pre-train to accommodate the continuing memorization of the corpus.
Outcome: The proposed method significantly outperforms existing pre-training generative retrieval baselines and performs well even in zero-shot settings.
PreSumm: Predicting Summarization Performance Without Summarizing (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in summarization models do not produce all documents in the same way, despite their inherent design principles and operational mechanisms.
Approach: They propose a task where a system predicts summarization performance based solely on the source document.
Outcome: The proposed task identifies documents that require manual summarization and improves dataset quality by filtering outliers and noisy documents.
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA (2025.emnlp-main)

Copied to clipboard

Challenge: Existing Question Answering systems are limited by noisy documents and flawed QA pairs.
Approach: They propose a high-quality subset of NarrativeQA focused on literary works . they identify and correct low-quality QA samples while removing extraneous text .
Outcome: The proposed subset of NarrativeQA is based on literary works.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations